Papers with multimodal cooking procedures
Predicting Implicit Arguments in Procedural Video Instructions (2025.acl-long)
Copied to clipboard
| Challenge: | Prior SRL benchmarks often miss implicit arguments, leading to incomplete understanding. |
| Approach: | They propose a dataset that necessitates inferring implicit and explicit arguments from contextual information in multimodal cooking procedures. |
| Outcome: | The proposed dataset achieves a 17% relative improvement in F1-score for what-implicit and a 14.7% improvement for where/with-implicative semantic roles over GPT-4o. |